Papers by Menno van Zaanen
Improving Machine Translation of Educational Content via Crowdsourcing (L18-1)
Copied to clipboard
Maximiliana Behnke, Antonio Valerio Miceli Barone, Rico Sennrich, Vilelmini Sosoni, Thanasis Naskos, Eirini Takoulidou, Maria Stasimioti, Menno van Zaanen, Sheila Castilho, Federico Gaspari, Panayota Georgakopoulou, Valia Kordoni, Markus Egg, Katia Lida Kermanidis
| Challenge: | Using crowdsourcing to train neural machine translation models is expensive and expensive . professional outsourcing of bilingual data is expensive if the translations are of a lower quality . |
| Approach: | They analyze the impact of crowdsourcing on the quality of in-domain training data . they use translations of MOOCs from English to eleven languages to fine-tune machine translation models . |
| Outcome: | The proposed method improves on general-domain training data and with pre-existing in-domain corpora. |
Translation Crowdsourcing: Creating a Multilingual Corpus of Online Educational Content (L18-1)
Copied to clipboard
Vilelmini Sosoni, Katia Lida Kermanidis, Maria Stasimioti, Thanasis Naskos, Eirini Takoulidou, Menno van Zaanen, Sheila Castilho, Panayota Georgakopoulou, Valia Kordoni, Markus Egg
| Challenge: | a large corpus of online content has been developed via large-scale crowdsourcing. |
| Approach: | They describe a multilingual corpus of online content that has been manually translated into 11 European and BRIC languages using the crowdsourcing platform. |
| Outcome: | The proposed corpus is a product of the EU-funded TraMOOC project and is used to train, tune and test machine translation engines. |
A Multilingual Wikified Data Set of Educational Material (L18-1)
Copied to clipboard
Iris Hendrickx, Eirini Takoulidou, Thanasis Naskos, Katia Lida Kermanidis, Vilelmini Sosoni, Hugo de Vos, Maria Stasimioti, Menno van Zaanen, Panayota Georgakopoulou, Valia Kordoni, Maja Popovic, Markus Egg, Antal van den Bosch
| Challenge: | a crowdsourcing effort to annotate and link parallel texts has been unsuccessful . a data set of parallel texts in eleven languages is presented . |
| Approach: | They present a wikified data set of English sentences linked to Wikipedia pages . they use crowdsourcing to annotate the texts and perform crowdsourcing for complex annotations . |
| Outcome: | The proposed data set is valuable as it constitutes a rich resource . it includes annotated data of English sentences linked to translations in eleven languages . |
A Process-oriented Dataset of Revisions during Writing (2020.lrec-1)
Copied to clipboard
| Challenge: | Revisions are defined as "changes at any point in the writing process" a dataset of 7,120 revisions was created to analyze revisions in writing . |
| Approach: | They use keystroke data and eye tracking data of 65 students to analyze revisions . they define revisions as "changes at any point in the writing process" |
| Outcome: | The proposed dataset includes 7,120 revisions from 65 students from different backgrounds . each type of revision can have a different effect on the written product or writing quality . |
Detecting Multiple Transitions in Literary Texts (2022.lrec-1)
Copied to clipboard
| Challenge: | Existing systems that can detect multiple transitions in texts are ineffective due to the large amount of texts available. |
| Approach: | They propose a system that can detect multiple transitions in literary texts . they extend existing system so it can detect transitions, and introduce multiple transition topics . |
| Outcome: | The proposed system outperforms the existing system on texts with known transitions and on single boundary texts. |